Micron Document
`:top
In `F33f`_`[mathematics`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Mathematics]`_`f, the `!matrix sign function`! is a `F33f`_`[matrix function`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Matrix_function]`_`f on `F33f`_`[square matrices`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Square_matrix]`_`f analogous to the complex `F33f`_`[sign function`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Sign_function]`_`f.`:cite-ref-0-1-0[`F5bf`_`[1`#cite-note-0-1]`_`f]

It was introduced by J.D. Roberts in 1971 as a tool for model reduction and for solving `F33f`_`[Lyapunov`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Lyapunov_equation]`_`f and Algebraic `F33f`_`[Riccati`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Algebraic_Riccati_equation]`_`f equation in a technical report of `F33f`_`[Cambridge University`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Cambridge_University]`_`f, which was later published in a journal in 1980.`:cite-ref-1-2-0[`F5bf`_`[2`#cite-note-1-2]`_`f]`:cite-ref-2-3-0[`F5bf`_`[3`#cite-note-2-3]`_`f]

>>Contents

• `F0af`_`[Definition`#definition]`_`f
• `F0af`_`[Properties`#properties]`_`f
• `F0af`_`[Computational methods`#computational-methods]`_`f
• `F0af`_`[Newton iteration`#newton-iteration]`_`f
• `F0af`_`[Newton–Schulz iteration`#newton-schulz-iteration]`_`f
• `F0af`_`[Applications`#applications]`_`f
• `F0af`_`[Solutions of Sylvester equations`#solutions-of-sylvester-equations]`_`f
• `F0af`_`[Solutions of algebraic Riccati equations`#solutions-of-algebraic-riccati-equations]`_`f
• `F0af`_`[Computations of matrix square-root`#computations-of-matrix-square-root]`_`f
• `F0af`_`[References`#references]`_`f

-─

>>Definition

The matrix sign function is a generalization of the complex `F33f`_`[signum function`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Sign_function]`_`f

csgn ⁡ ⁡ ( z ) = { 1 if R e ( z ) > 0 , − − 1 if R e ( z ) < 0 , {\\displaystyle \\operatorname {csgn} (z)={\\begin{cases}1&{\\text{if }}\\mathrm {Re} (z)>0,\\\\-1&{\\text{if }}\\mathrm {Re} (z)<0,\\end{cases}}}

to the matrix valued analogue csgn ⁡ ⁡ ( A ) {\\displaystyle \\operatorname {csgn} (A)} . Although the sign function is not `F33f`_`[analytic`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Analytic_function]`_`f, the `F33f`_`[matrix function`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Matrix_function]`_`f is well defined for all matrices that have no `F33f`_`[eigenvalue`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Eigenvalues_and_eigenvectors]`_`f on the `F33f`_`[imaginary axis`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Imaginary_axis]`_`f, see for example the `F33f`_`[Jordan-form-based definition`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Analytic_function_of_a_matrix]`_`f (where the derivatives are all zero).

>>Properties

`!Theorem:`! Let A ∈ ∈ C n × × n {\\displaystyle A\\in \\mathbb {C} ^{n\\times n}} , then csgn ⁡ ⁡ ( A ) 2 = I {\\displaystyle \\operatorname {csgn} (A)^{2}=I} .`:cite-ref-0-1-1[`F5bf`_`[1`#cite-note-0-1]`_`f]

`!Theorem:`! Let A ∈ ∈ C n × × n {\\displaystyle A\\in \\mathbb {C} ^{n\\times n}} , then csgn ⁡ ⁡ ( A ) {\\displaystyle \\operatorname {csgn} (A)} is `F33f`_`[diagonalizable`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Diagonalizable_matrix]`_`f and has eigenvalues that are ± ± 1 {\\displaystyle \\pm 1} .`:cite-ref-0-1-2[`F5bf`_`[1`#cite-note-0-1]`_`f]

`!Theorem:`! Let A ∈ ∈ C n × × n {\\displaystyle A\\in \\mathbb {C} ^{n\\times n}} , then ( I + csgn ⁡ ⁡ ( A ) ) / 2 {\\displaystyle (I+\\operatorname {csgn} (A))/2} is a `F33f`_`[projector`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Projection_(linear_algebra)]`_`f onto the `F33f`_`[invariant subspace`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Invariant_subspace]`_`f associated with the eigenvalues in the `F33f`_`[right-half plane`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Right_half-plane]`_`f, and analogously for ( I − − csgn ⁡ ⁡ ( A ) ) / 2 {\\displaystyle (I-\\operatorname {csgn} (A))/2} and the left-half plane.`:cite-ref-0-1-3[`F5bf`_`[1`#cite-note-0-1]`_`f]

`!Theorem:`! Let A ∈ ∈ C n × × n {\\displaystyle A\\in \\mathbb {C} ^{n\\times n}} , and A = P [ J + 0 0 J − − ] P − − 1 {\\displaystyle A=P{\\begin{bmatrix}J_{+}&0\\\\0&J_{-}\\end{bmatrix}}P^{-1}} be a `F33f`_`[Jordan decomposition`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Jordan_normal_form]`_`f such that J + {\\displaystyle J_{+}} corresponds to eigenvalues with positive real part and J − − {\\displaystyle J_{-}} to eigenvalue with negative real part. Then csgn ⁡ ⁡ ( A ) = P [ I + 0 0 − − I − − ] P − − 1 {\\displaystyle \\operatorname {csgn} (A)=P{\\begin{bmatrix}I_{+}&0\\\\0&-I_{-}\\end{bmatrix}}P^{-1}} , where I + {\\displaystyle I_{+}} and I − − {\\displaystyle I_{-}} are `F33f`_`[identity matrices`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Identity_matrix]`_`f of sizes corresponding to J + {\\displaystyle J_{+}} and J − − {\\displaystyle J_{-}} , respectively.`:cite-ref-0-1-4[`F5bf`_`[1`#cite-note-0-1]`_`f]

>>Computational methods

The function can be computed with generic methods for `F33f`_`[matrix functions`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Matrix_function]`_`f, but there are also specialized methods.

>>>Newton iteration

The `F33f`_`[Newton iteration`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Newton's_method]`_`f can be derived by observing that csgn ⁡ ⁡ ( x ) = x 2 / x {\\displaystyle \\operatorname {csgn} (x)={\\sqrt {x^{2}}}/x} , which in terms of matrices can be written as csgn ⁡ ⁡ ( A ) = A − − 1 A 2 {\\displaystyle \\operatorname {csgn} (A)=A^{-1}{\\sqrt {A^{2}}}} , where we use the `F33f`_`[matrix square root`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Matrix_square_root]`_`f. If we apply the `F33f`_`[Babylonian method`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Square_root_of_a_matrix]`_`f to compute the square root of the matrix A 2 {\\displaystyle A^{2}} , that is, the iteration X k + 1 = 1 2 ( X k + A 2 X k − − 1 ) {\\textstyle X_{k+1}={\\frac {1}{2}}\\left(X_{k}+A^{2}X_{k}^{-1}\\right)} , and define the new iterate Z k = A − − 1 X k {\\displaystyle Z_{k}=A^{-1}X_{k}} , we arrive at the iteration

Z k + 1 = 1 2 ( Z k + Z k − − 1 ) {\\displaystyle Z_{k+1}={\\frac {1}{2}}\\left(Z_{k}+Z_{k}^{-1}\\right)} ,

where typically Z 0 = A {\\displaystyle Z_{0}=A} . Convergence is global, and locally it is quadratic.`:cite-ref-0-1-5[`F5bf`_`[1`#cite-note-0-1]`_`f]`:cite-ref-1-2-1[`F5bf`_`[2`#cite-note-1-2]`_`f]

The Newton iteration uses the explicit inverse of the iterates Z k {\\displaystyle Z_{k}} .

>>>Newton–Schulz iteration

To avoid the need of an explicit inverse used in the Newton iteration, the inverse can be approximated with one step of the `F33f`_`[Newton iteration for the inverse`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Newton's_method]`_`f, Z k − − 1 ≈ ≈ Z k ( 2 I − − Z k 2 ) {\\displaystyle Z_{k}^{-1}\\approx Z_{k}\\left(2I-Z_{k}^{2}\\right)} , derived by Schulz(de) in 1933.`:cite-ref-4[`F5bf`_`[4`#cite-note-4]`_`f] Substituting this approximation into the previous method, the new method becomes

Z k + 1 = 1 2 Z k ( 3 I − − Z k 2 ) {\\displaystyle Z_{k+1}={\\frac {1}{2}}Z_{k}\\left(3I-Z_{k}^{2}\\right)} .

Convergence is (still) quadratic, but only local (guaranteed for ‖ ‖ I − − A 2 ‖ ‖ < 1 {\\displaystyle \\|I-A^{2}\\|<1} ).`:cite-ref-0-1-6[`F5bf`_`[1`#cite-note-0-1]`_`f]

>>Applications

>>>Solutions of Sylvester equations

`!Theorem:`:cite-ref-1-2-2[`F5bf`_`[2`#cite-note-1-2]`_`f]`:cite-ref-2-3-1[`F5bf`_`[3`#cite-note-2-3]`_`f]`! Let A , B , C ∈ ∈ R n × × n {\\displaystyle A,B,C\\in \\mathbb {R} ^{n\\times n}} and assume that A {\\displaystyle A} and B {\\displaystyle B} are `F33f`_`[stable`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Stable_matrix]`_`f, then the unique solution to the `F33f`_`[Sylvester equation`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Sylvester_equation]`_`f, A X + X B = C {\\displaystyle AX+XB=C} , is given by X {\\displaystyle X} such that

[ − − I 2 X 0 I ] = csgn ⁡ ⁡ ( [ A − − C 0 − − B ] ) . {\\displaystyle {\\begin{bmatrix}-I&2X\\\\0&I\\end{bmatrix}}=\\operatorname {csgn} \\left({\\begin{bmatrix}A&-C\\\\0&-B\\end{bmatrix}}\\right).}

`*Proof sketch:`* The result follows from the similarity transform

[ A − − C 0 − − B ] = [ I X 0 I ] [ A 0 0 − − B ] [ I X 0 I ] − − 1 , {\\displaystyle {\\begin{bmatrix}A&-C\\\\0&-B\\end{bmatrix}}={\\begin{bmatrix}I&X\\\\0&I\\end{bmatrix}}{\\begin{bmatrix}A&0\\\\0&-B\\end{bmatrix}}{\\begin{bmatrix}I&X\\\\0&I\\end{bmatrix}}^{-1},}

since

csgn ⁡ ⁡ ( [ A − − C 0 − − B ] ) = [ I X 0 I ] [ I 0 0 − − I ] [ I − − X 0 I ] , {\\displaystyle \\operatorname {csgn} \\left({\\begin{bmatrix}A&-C\\\\0&-B\\end{bmatrix}}\\right)={\\begin{bmatrix}I&X\\\\0&I\\end{bmatrix}}{\\begin{bmatrix}I&0\\\\0&-I\\end{bmatrix}}{\\begin{bmatrix}I&-X\\\\0&I\\end{bmatrix}},}

due to the stability of A {\\displaystyle A} and B {\\displaystyle B} .

The theorem is, naturally, also applicable to the `F33f`_`[Lyapunov equation`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Lyapunov_equation]`_`f. However, due to the structure the Newton iteration simplifies to only involving inverses of A {\\displaystyle A} and A T {\\displaystyle A^{T}} .

>>>Solutions of algebraic Riccati equations

There is a similar result applicable to the `F33f`_`[algebraic Riccati equation`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Algebraic_Riccati_equation]`_`f, A H P + P A − − P F P + Q = 0 {\\displaystyle A^{H}P+PA-PFP+Q=0} .`:cite-ref-0-1-7[`F5bf`_`[1`#cite-note-0-1]`_`f]`:cite-ref-1-2-3[`F5bf`_`[2`#cite-note-1-2]`_`f] Define V , W ∈ ∈ C 2 n × × n {\\displaystyle V,W\\in \\mathbb {C} ^{2n\\times n}} as

[ V W ] = csgn ⁡ ⁡ ( [ A H Q F − − A ] ) − − [ I 0 0 I ] . {\\displaystyle {\\begin{bmatrix}V&W\\end{bmatrix}}=\\operatorname {csgn} \\left({\\begin{bmatrix}A^{H}&Q\\\\F&-A\\end{bmatrix}}\\right)-{\\begin{bmatrix}I&0\\\\0&I\\end{bmatrix}}.}

Under the assumption that F , Q ∈ ∈ C n × × n {\\displaystyle F,Q\\in \\mathbb {C} ^{n\\times n}} are `F33f`_`[Hermitian`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Hermitian_matrix]`_`f and there exists a unique stabilizing solution, in the sense that A − − F P {\\displaystyle A-FP} is `F33f`_`[stable`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Stable_matrix]`_`f, that solution is given by the `F33f`_`[over-determined`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Overdetermined_system]`_`f, but `F33f`_`[consistent`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Consistent_and_inconsistent_equations]`_`f, linear system

V P = − − W . {\\displaystyle VP=-W.}

`*Proof sketch:`* The similarity transform

[ A H Q F − − A ] = [ P − − I I 0 ] [ − − ( A − − F P ) − − F 0 ( A − − F P ) ] [ P − − I I 0 ] − − 1 , {\\displaystyle {\\begin{bmatrix}A^{H}&Q\\\\F&-A\\end{bmatrix}}={\\begin{bmatrix}P&-I\\\\I&0\\end{bmatrix}}{\\begin{bmatrix}-(A-FP)&-F\\\\0&(A-FP)\\end{bmatrix}}{\\begin{bmatrix}P&-I\\\\I&0\\end{bmatrix}}^{-1},}

and the stability of A − − F P {\\displaystyle A-FP} implies that

( csgn ⁡ ⁡ ( [ A H Q F − − A ] ) − − [ I 0 0 I ] ) [ X − − I I 0 ] = [ X − − I I 0 ] [ 0 Y 0 − − 2 I ] , {\\displaystyle \\left(\\operatorname {csgn} \\left({\\begin{bmatrix}A^{H}&Q\\\\F&-A\\end{bmatrix}}\\right)-{\\begin{bmatrix}I&0\\\\0&I\\end{bmatrix}}\\right){\\begin{bmatrix}X&-I\\\\I&0\\end{bmatrix}}={\\begin{bmatrix}X&-I\\\\I&0\\end{bmatrix}}{\\begin{bmatrix}0&Y\\\\0&-2I\\end{bmatrix}},}

for some matrix Y ∈ ∈ C n × × n {\\displaystyle Y\\in \\mathbb {C} ^{n\\times n}} .

>>>Computations of matrix square-root

The `F33f`_`[Denman–Beavers iteration`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Square_root_of_a_matrix]`_`f for the square root of a matrix can be derived from the Newton iteration for the matrix sign function by noticing that A − − P I P = 0 {\\displaystyle A-PIP=0} is a degenerate `F33f`_`[algebraic Riccati equation`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Algebraic_Riccati_equation]`_`f`:cite-ref-2-3-2[`F5bf`_`[3`#cite-note-2-3]`_`f] and by definition a solution P {\\displaystyle P} is the square root of A {\\displaystyle A} .

>>References

`:cite-note-0-1`!1.`! `F0af`_`[↑`#cite-ref-0-1-0]`_`f `:citerefhigham2008`a`F33f`_`[Higham, Nicholas J.`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Nicholas_Higham]`_`f (2008). `*Functions of matrices : theory and computation`*. Society for Industrial and Applied Mathematics. Philadelphia, Pa.: Society for Industrial and Applied Mathematics (SIAM, 3600 Market Street, Floor 6, Philadelphia, PA 19104). `F33f`_`[ISBN`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=ISBN_(identifier)]`_`f 978-0-89871-777-8. `F33f`_`[OCLC`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=OCLC_(identifier)]`_`f 693957820.
`:cite-note-1-2`!2.`! `F0af`_`[↑`#cite-ref-1-2-0]`_`f `:citerefroberts1980`aRoberts, J. D. (October 1980). "Linear model reduction and solution of the algebraic Riccati equation by use of the sign function". `*International Journal of Control`*. `!32`! (4): 677–687. `F33f`_`[doi`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Doi_(identifier)]`_`f:10.1080/00207178008922881. `F33f`_`[ISSN`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=ISSN_(identifier)]`_`f 0020-7179.
`:cite-note-2-3`!3.`! `F0af`_`[↑`#cite-ref-2-3-0]`_`f `:citerefdenmanbeavers1976`aDenman, Eugene D.; Beavers, Alex N. (1976). "The matrix sign function and computations in systems". `*Applied Mathematics and Computation`*. `!2`! (1): 63–94. `F33f`_`[doi`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Doi_(identifier)]`_`f:10.1016/0096-3003(76)90020-5. `F33f`_`[ISSN`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=ISSN_(identifier)]`_`f 0096-3003.
`:cite-note-4`!4.`! `F0af`_`[↑`#cite-ref-4]`_`f `:citerefschulz1933`aSchulz, Günther (1933). "Iterative Berechung der reziproken Matrix". `*ZAMM - Journal of Applied Mathematics and Mechanics / Zeitschrift für Angewandte Mathematik und Mechanik`*. `!13`! (1): 57–59. `F33f`_`[Bibcode`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Bibcode_(identifier)]`_`f:1933ZaMM...13...57S. `F33f`_`[doi`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=Doi_(identifier)]`_`f:10.1002/zamm.19330130111. `F33f`_`[ISSN`:/page/wikibook/entry.mu`zim=wikipedia_en_all_nopic_2025-08.zim|entry_path=ISSN_(identifier)]`_`f 1521-4001.

`c`F0af`_`[↑ Back to top`#top]`_`f`a